Back

BioData Mining

Springer Science and Business Media LLC

Preprints posted in the last 7 days, ranked by how well they match BioData Mining's content profile, based on 22 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Automatic bioinformatic software named entity recognition from literature

Xuan, H.; Pasupuleti, R.; Liu, B.; Sun, H.; Zhang, J.; Yao, Z.; Zhong, C.

2026-09-01 bioinformatics 10.64898/2026.08.26.731133 medRxiv
Top 0.1%
6.2%
Show abstract

Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.

2
Electronic health data exploring cardiorespiratory responses of transfusions in preterm infants: An international multicenter cohort study

Honore, A.; Rech, T.; Scrivens, A.; Binotto, I.; Zandvoort, C. S.; van der Staaij, H.; Peck, M.; Zivanovic, S.; Stanworth, S. J.; Hartley, C.; Dame, C.; Deschmann, E.; the Neonatal Transfusion Network,

2026-09-03 pediatrics 10.64898/2026.09.01.26361418 medRxiv
Top 0.1%
5.6%
Show abstract

Background and Objectives: Preterm infants are commonly transfused, yet direct cardiorespiratory effects of red blood cell (RBC) transfusions remain poorly understood. We explored the feasibility of using multicentre electronic health data (EHD) to study such cardiorespiratory responses. Methods: Highly granular routine EHD were collected from preterm infants born <32 weeks gestational age at three European centres. Heart rate, oxygen saturation, and respiratory rate were evaluated 12 hours before and after the RBC transfusion. Results: A total of 321 transfusions in 164 infants were analysed. Overall, there was no significant change in the rate of bradycardia and apnoea following transfusion. Cardiorespiratory parameters varied substantially between infants; e.g. 20% of transfusions were associated with an unexpected, significant increase in heart rate. Respiratory rate and oxygen saturation exhibited similarly heterogenous patterns following transfusion. In sub-group analysis, the proportion of transfusions with increased heart rate was significantly higher within the first two weeks than later (32% vs 13%, p=0.0019). Conclusions: Multicentre EHD extraction allows to identify otherwise masked short-term effects of RBC transfusions on cardiorespiratory parameters, possibly indicating cardiac or pulmonary overload. Such effects may vary with adaptation to anaemia. Analysing EHD may ultimately enable personalized transfusion practice.

3
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.1%
3.5%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

4
An M-learner approach for heterogeneous mediation analysis with high-dimensional omics mediators

Li, X.; Wei, P.

2026-09-01 bioinformatics 10.64898/2026.08.25.747106 medRxiv
Top 0.1%
3.1%
Show abstract

Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.

5
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.2%
1.9%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

6
Every Cure Knowledge Graph: A Unified Biomedical Knowledge Graph for Drug Repurposing

Kaniewski, P.; Carter, E. K.; Rhodes, D.; Lim, E. M.; Li, J.; Vergine, J.; Matentzoglu, N.; Schaper, K.; Reilly, J.; Sundar, S.; Vijnck, L.; Sharp, E.; Alfonso, N.; Ford, A.; Stepanenko, A.; Hempstead, C.; Brokmeier, P.; Bizon, C.; Tropsha, A.; Haendel, M. A.; Fajgenbaum, D. C.; Lancashire, L.

2026-08-31 bioinformatics 10.64898/2026.08.26.747253 medRxiv
Top 0.3%
1.4%
Show abstract

Identifying causal connections between existing drugs and mechanistic profiles of diseases is a foundational step for effective drug repurposing. Although knowledge graphs (KGs) are highly suited for consolidating biomedical databases and tracking these connections, a single biomedical KG is constrained by its ingestion pipeline and knowledge sources. While different biomedical KGs could be complementary if combined, efforts to combine them into a unified and more comprehensive KG are hindered by lack of interoperability and poor provenance. To address those issues, we present EC-KG, a Biolink Model-compatible KG for computational drug repurposing. EC-KG is an interoperable, provenance-first KG which integrates RTX-KG2, ROBOKOP, and PrimeKG at the network-level, encapsulating over 7 million nodes and 81 million edges from 95 primary data sources. EC-KG has improved coverage of core biomedical entities such as drugs, targets, and diseases relevant to drug repurposing vs source graphs, and captures complex biomedical mechanisms within its topology. We demonstrate that the network unification in EC-KG leads to emergence of novel, mechanistically relevant pathways which are disconnected in the underlying constituent networks and show its applications in method development, benchmarking and predictive drug repurposing applications. EC-KG has already been successfully used in drug repurposing research to surface Botulinum Toxin A as a candidate to treat Major Depressive Disorder, as well as to validate repurposing of Lenalidomide and Dexamethasone for a subgroup of patients with Rosai-Dorfman Disease.

7
Constructing microbiome co-occurrence networks with confidence: A conditional, nonparametric, inference-based approach

Song, H.; Xiang, Y.; Liu, H.; Ling, W.; Plantinga, A. M.; Srinivasan, S.; Dun, Y.; Zhao, N.; Sun, S.; Engel, S. M.; Simon, N.; Wu, M. C.

2026-09-01 bioinformatics 10.64898/2026.08.27.747483 medRxiv
Top 0.3%
1.3%
Show abstract

Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.

8
Machine learning analysis of Autism phenotype data supports a four-dimensional continuum with three overlapping subtypes

Quigley, H.; Gardiner, B.; McDaid, L.; O'Donnell, C.

2026-08-31 psychiatry and clinical psychology 10.64898/2026.08.27.26361561 medRxiv
Top 0.4%
1.1%
Show abstract

Autism Spectrum Disorder (ASD) is a heterogeneous neurodevelopmental condition defined by differences in social communication and restricted, repetitive behaviours. As diagnostic criteria have broadened, ASD is now recognised across a wider range of individuals, raising key questions about its structure: does ASD have discrete sub-types, or is it better conceptualised as a continuous, possibly multidimensional, condition? We aim to explore whether a multidimensional continuum model more accurately captures the variability within ASD. We analysed a large SPARK phenotypic dataset of medical history and diagnostic surveys (background history, SCQ, RBS-R; n=36,710 individuals). We apply and compare two traditional statistical approaches, Factor Analysis and Gaussian Mixture Models, with a modern machine learning technique, the Variational Autoencoder (VAE). VAEs reconstructed unseen test data with ~4-fold better accuracy than Factor Analysis, and ~8-fold better accuracy than Gaussian Mixture Models. We identified four stable latent factors across 100 independently trained VAEs. These four dimensions provide an individual behavioural profile that can be visualized using radar-plots, offering a compact way to compare profiles at the person level. Through further analysis, we found evidence for 3 overlapping clusters or subtypes of ASD identified within the 4D latent space. This work aims to inform new ways of modelling ASD using a VAE that will be able to discern between a continuum or a clustered output and that go beyond binary diagnosis, instead reflecting the complex range of trait profiles, with implications for personalised diagnosis and intervention.

9
Novel Large Language Model-Based Detection of Echocardiographic Markers of Right Ventricular Dysfunction

Ekambarapu, L.; Pendyal, A.; Lin, A.; Alwakeel, M.; Rajaratnam, A.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361456 medRxiv
Top 0.4%
1.1%
Show abstract

Background: Unstructured biomedical data, such as echocardiography reports, are rich in information but time consuming to analyze at scale. Rule-based, regular expression-driven terminology mapping can only extract individual variables while large language models (LLMs) offer scalable and clinically meaningful interpretations of heterogeneous disease processes. Right ventricular dysfunction (RVD) is an example of a multifactorial disease state in which key structural and physiologic features are captured both narratively and in structured fields, making it an ideal test case for evaluating whether LLMs can recover complex phenotypes that rules based methods routinely miss. Purpose: To compare an LLM-based extraction method to a conventional rules-based schema for identifying and phenotyping echocardiographic features associated with RVD in a large TTE dataset. Methods: MIMIC-III NOTE2NUM echocardiography reports (n = 45,794) were analyzed using GPT-4o-based LLM extraction deployed within a secure health system enclave and were benchmarked against echocardiographic measurements defined in the MIMIC-III dictionary schema. In MIMIC-III, PH was recorded qualitatively (mild/moderate/severe) based on tricuspid regurgitant (TR) jet velocity and then re-coded as present vs. absent. LLM based extraction defined RVD as (1) RV structural abnormality (>= 1 of hypertrophy, dilation, or wall hypo-/akinesis) or (2) RV pressure/volume overload (>= 2 of the following: estimated right atrial pressure > 8 mmHg, TR jet velocity > 2.8 m/s, fractional area change < 35%, tricuspid annular planar systolic excursion < 17 mm, S' < 9.5 cm/s, or E/e' > 14), with PH defined as estimated pulmonary artery systolic pressure > 35 mmHg or qualitative documentation of PH. Results: LLM extraction identified PH in 15,394 (33.6%), RV pressure/volume overload in 14,449 (31.6%), and RV structural abnormalities in 11,955 (26.1%). Co-occurrence was common: overload + structural changes in 9,380 (20.5%), overload + PH in 9,756 (21.3%), structural changes + PH in 6,183 (13.5%), and all three in 5,620 (12.3%). Using the MIMIC-III dictionary schema, PH prevalence was similar (15,371; 33.6%), but RV overload fields were captured less often (pressure overload 1,357 [3.0%], volume overload 1,128 [2.5%], pressure + volume overload 1,093 [2.4%]; any overload field 3,578 [7.8%]), and RV pressure/volume overload with PH was identified in only 731 (1.6%). Conclusions: LLM-based extraction outperforms rules-based schemas for identifying complex disease states not defined by any single variable. By synthesizing multifactorial signals, LLMs can phenotype RVD with higher fidelity and support population-level assessment. Further validation using multimodality imaging, invasive hemodynamics, and clinical outcome data is needed.

10
Rising rate of non-receipt of vitamin K prophylaxis for newborns, January 2019 - June 2026

Masters, N. B.; Farrar, K. G.; Holler, E.; Lancaster, J. M.

2026-09-02 pediatrics 10.64898/2026.08.31.26361837 medRxiv
Top 0.5%
1.0%
Show abstract

Background: Vitamin K prophylaxis is universally recommended for newborns to prevent life threatening vitamin K deficiency bleeding. Although not on the immunization schedule, vitamin K prophylaxis is often coadministered with hepatitis B birth dose and erythromycin ophthalmic ointment, and rising hesitancy around vaccines/preventive care may spill over into vitamin K administration. Methods: We conducted a retrospective cohort study using Truveta electronic health record data with linked mother-child dyads. Live births to mothers aged 15-49 from January 1, 2019 through June 30, 2026 were included. Vitamin K administration was defined as documentation on the birth date or following day. Logistic regression assessed sociodemographic predictors of non-receipt, and interrupted time series analysis evaluated changes after January 2026. Results: Among 1,026,375 infants, 995,628 (96.97%) had documented vitamin K administration. Non-receipt increased from an average of 2.1% during 2019-2022 to 4.3% in 2025 and 6.1% in 2026, reaching 8.10% in June 2026. Older maternal age, non-Hispanic or Latino ethnicity, Medicaid or unknown insurance, and year of delivery were associated with greater odds of non-receipt. After January 2026, there was no immediate step change, but the odds of vitamin K receipt declined an additional 10% per month (OR: 0.90; 95% CI, 0.88-0.91). Conclusions: Vitamin K non-receipt increased over the study period and accelerated after January 2026. Because vitamin K recommendations were not changed by the January vaccine schedule, this association may reflect broader impacts to confidence in newborn preventive care. Future studies should examine causal mechanisms, parental decision-making, and associated clinical outcomes.

11
Dynamic Clinical States and Transitions During the First 72 Hours of Intensive Care After Acute Stroke

LEI, P.; XU, Y.; ZHANG, Y.

2026-09-01 intensive care and critical care medicine 10.64898/2026.08.30.26361738 medRxiv
Top 0.5%
1.0%
Show abstract

Background: The condition of a patient with acute stroke often changes within hours of ICU admission. Prognostic work here targets fixed endpoints predicted from admission data, and trajectory phenotyping assigns one label per patient. We used longitudinal ICU data to identify interpretable dynamic clinical states, characterize transitions between them, and relate the current state to later events. Methods: Retrospective cohort study of 6368 adults with acute stroke in MIMIC IV v3.1. The first 72 h were divided into twelve 6-hour windows, and a hidden Markov model was fitted to 21 neurological, physiological and organ support variables. State number was chosen against criteria fixed before fitting: statistical fit, restart stability, state occupancy and clinical interpretability. Generalized estimating equations related the current state to new mechanical ventilation and vasopressor use within 12 h, and to ICU death within 72 h. Eleven sensitivity analyses assessed the robustness of the state solution. Results: Four states were selected: neurologically preserved-low support, neurological impairment low support, impairment renal dysfunction and impairment-respiratory support (63.3%, 7.8%, 11.8% and 17.1% of windows). Within 72 h, 40.3% of patients changed state at least once, and transitions ran in both directions rather than along a single severity gradient. States were identified without outcome data, yet ICU mortality by last state ranged from 2.9% to 43.9%. Adjusted for age, sex, subtype and Charlson index, the current state remained associated with organ-support escalation and death. State prevalence differed by at most 1.1 percentage points between training and test sets, and 10 of 11 sensitivity analyses gave a stable four-state solution (ARI 0.754 0.955). Conclusions: The early ICU course of acute stroke can be represented as movement among a small number of clinically interpretable states. The representation was reproducible in a held out set and across admission eras, but requires validation in an independent database before any clinical use.

12
Mural-VISTA: A tool for mural cell-vessel interaction assessment and multiscale single-cell topo-morphological analysis

Zeng, H.; Hu, M.; Phng, L.-K.; Matsunaga, Y. T.

2026-09-01 bioinformatics 10.64898/2026.08.27.747487 medRxiv
Top 0.6%
0.9%
Show abstract

Three-dimensional (3D) mural cell morphology is heterogeneous and coupled to vessel geometry, however, measurements from two-dimensional (2D) maximum intensity projections (MIP) obscure overlapping processes and cell-vessel contacts. Accordingly, we developed Mural-VISTA, a semi-automated Python workflow for mural cell-vessel interaction and single-cell topo-morphology analysis of reconstructed surface meshes. This workflow integrates mesh pretreatment, interactive centerline extraction, hierarchical segmentation of cell soma, main axis and secondary processes (branches), and extraction of 36 multiscale (cell process segment level, process level, and whole cell level) topo-morphological and vessel-referenced metrics. Mural-VISTA identified morphological changes in pericytes and vascular smooth muscle cells (vSMCs) with altered RhoA activity. Constitutive active RhoA (RhoA CA) over-expression reduced branch complexity and increased process alignment in both cell types, while increased whole-cell and branch solidity only in vSMCs. Dominant negative RhoA (RhoA DN) over-expression increased branch abundance and reduced branch solidity in pericytes but not vSMCs, suggesting cell-type specific effect of reduced RhoA activity. In conclusion, Mural-VISTA enables quantitative 3D profiling of mural cell architecture and its spatial relationship with the vessel.

13
RedFuMOS: A novel approach for multi-omics and clinical data-driven patient stratification

De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.

2026-08-31 health informatics 10.64898/2026.08.26.26361415 medRxiv
Top 0.7%
0.8%
Show abstract

Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.

14
Prospective In-silico Simulation of the VESALIUS-CV Trial Using Biomedical Knowledge Graph and Real-World Data-Driven AI Modeling

Perlman, A.; Goldstein, N.; Goldman, M.; Shapiro, M.; Barash, E.; Bar, A.; Raveh, T.; Tordjman, E.; Schussheim, H.; Dormont, F.; Matalon, O.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361436 medRxiv
Top 0.9%
0.6%
Show abstract

Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulation using real-world data (RWD) has emerged as a potential tool to support earlier decision-making; however, evidence of prospective predictive validity, generated prior to trial result disclosure, remains limited. Methods. We applied a semi-mechanistic machine learning framework integrating real-world patient data with biologically informed drug representations to prospectively simulate the VESALIUS-CV trial evaluating evolocumab versus placebo. The simulation model was trained on a combination of patient-level real-world data and a drug-centric knowledge graph and validated for both patient-level and trial-level retrospective predictive performance. The model was then used to simulate VESALIUS-CV before public disclosure of trial results, using a locked model and prespecified eligibility criteria and primary endpoint aligned with the clinical protocol. A patient-level time-to-event model was used to generate virtual trial arms, from which cumulative incidence curves, hazard ratios, confidence intervals, and p-values for major adverse cardiovascular events (MACE) were estimated. Results. In retrospective validation, the model demonstrated strong patient-level discrimination, with time-dependent ROC-AUC values ranging from 0.80 to 0.90 across follow-up horizons. For trial-level validation, 22 randomized cardiovascular-outcomes trials were simulated, and hazard ratios for 3-point MACE across 24 between-arm comparisons showed consistent directional agreement and quantitative correlation with published results such that the model accurately predicted trial success, achieving an F1 score of 0.83, with precision of 0.79 and sensitivity of 0.89. In a fully prospective application, the simulation predicted a statistically significant reduction in 3-point MACE with evolocumab versus placebo, estimating a hazard ratio of 0.78 (95% CI, 0.70-0.87) at 54 months. These predictions were consistent with the subsequently reported VESALIUS-CV results, which demonstrated a hazard ratio of 0.75 (95% CI, 0.65-0.86) at 55 months of median follow-up. Conclusions. In a fully prospective setting, a RWD-driven, AI-based simulation accurately predicted the direction, magnitude, and temporal dynamics of treatment effects observed in the VESALIUS-CV trial. These results demonstrate that in-silico trial simulation can anticipate clinical outcomes in the prospective setting, supporting its use as a complementary tool for early decision-making, trial design optimization, and de-risking in cardiovascular drug development.

15
Sphingolipid metabolism-related genes as key regulatory hubs in white smoke inhalation induced lung injury

Meng, F.; Xin, H.; Li, R. R.

2026-09-01 bioinformatics 10.64898/2026.08.26.747407 medRxiv
Top 1.0%
0.5%
Show abstract

Objective White smoke inhalation injury (WSI) causes severe acute lung damage with no specific therapy currently available. Sphingolipid metabolism is implicated in pulmonary inflammation, but its transcriptional regulatory landscape in WSI remains unexplored. This study aimed to identify key sphingolipid metabolism related genes and evaluate their regulatory roles and therapeutic potential in WSI. Methods We established a rat model of WSI and performed integrated bulk RNA sequencing, weighted gene coexpression network analysis (WGCNA), and single-cell RNA sequencing (scRNAseq) to screen for differentially expressed sphingolipid metabolism-related genes (DESRGs). Protein-protein interaction (PPI) network with four centrality algorithms was used to prioritize hub genes. In silico gene knockout and molecular docking were conducted to assess regulatory functions and identify potential drug candidates. Results We identified 22 DESRGs that were predominantly enriched in DNA replication and cell cycle pathways rather than canonical sphingolipid metabolic processes. PPI consensus prioritized three hub genes--Top2a, Ttk, and Ccna2--with Top2a exhibiting the highest expression in epithelial cells and significant downregulation after smoke exposure. ScRNAseq revealed immune cell infiltration and epithelial differentiation trajectories. Virtual knockout showed that Top2a depletion affected the largest transcriptomic fraction (~0.4%) and was enriched in lysosome biogenesis, innate immunity, phagocytosis, and lipid catabolism. Molecular docking identified thalidomide as a high affinity ligand for Top2a (Vina score: -8.5 kcal/mol). Conclusion Our multiomics integrative framework identifies Top2a as a central regulatory hub linking sphingolipid associated inflammation to epithelial responses in WSI, and nominates thalidomide as a potential drug repurposing candidate. These findings provide prioritized targets for future translational investigation.

16
Network-based meta-analysis maps stage-dependent molecular programs in MASLD through MASLD-META NETWORK application

Kumak, E.; Darde, T.; Konu, O.

2026-08-31 bioinformatics 10.64898/2026.08.26.747338 medRxiv
Top 1.0%
0.5%
Show abstract

Metabolic dysfunction-associated steatotic liver disease (MASLD), the leading cause of chronic liver pathologies worldwide, represents a growing clinical burden. Its diagnosis remains reliant on liver biopsy that limits early detection and the ability to capture molecular changes across disease progression. A systematic understanding of stage-dependent gene expression changes is essential to identify biomarkers and effectively characterize disease mechanisms. Therefore recent studies provided databases for searching genes as well as prediction of multi-gene signatures for disease progression. However, there is still a need for interactive and comprehensive meta-analysis of datasets of MASLD patients with available histological metadata. Herein, we performed a meta-analysis of RNA-seq datasets using NAFLD Activity Score (NAS; n = 897) and fibrosis stage (n = 856) upon conducting pairwise comparisons across histological stages and identified differentially expressed genes associated with disease progression. Most importantly, we provide our findings via a dedicated web server, the MASLD-META NETWORK (https://masld.scilicium.com), enabling users to interactively explore meta-analysis results across diverse network modalities. In addition, we characterized gene expression dynamics across increasing disease stages to identify consistent progression-associated pathways using Louvain clustering. Network-based parameters such as centrality in combination with meta-analysis scores further highlighted central genes and pathways implicated in disease mechanisms. Accordingly, MASLD-META NETWORK enabled an integrative reassessment of recently published gene signatures, identifying COL1A1, COL3A1, THBS2, FBLN5, and PDGFA as the most central genes, and SULF2, MMP14, IL32, GPNMB, and COL3A1 as candidate markers of earlier transcriptional alterations. Network analysis of MASLD associated biological modules further identified LAMA2 and LAMA3 as previously unrecognized central candidate targets.

17
Autotaxin Inhibition Ameliorates HFpEF Phenotype By Reducing LPA-Mediated Systemic Inflammation And Cardiac Remodeling

Chaudhary, R.; Robbins, A.; Singh, A. P.; Shabani, P.; Luther, T. K.; Alzamrooni, A.; Lopez, R.; Maheshwari, T.; Collins, N.; Hummel, S.; Abdel-Latif, A.

2026-08-30 immunology 10.64898/2026.08.26.747366 medRxiv
Top 1%
0.5%
Show abstract

Background: HFpEF accounts for roughly half of heart failure admissions and lacks disease-modifying therapy. Autotaxin (ENPP2) generates lysophosphatidic acid (LPA), a profibrotic and pro-inflammatory bioactive lipid. Whether circulating lysophospholipid metabolism is altered in HFpEF, and whether autotaxin inhibition modifies an established experimental HFpEF phenotype, is untested. Methods: Plasma from patients with HFpEF (n=210) and non-heart-failure comparators (n=27) underwent untargeted and LPA-targeted mass spectrometry and a nine-analyte multiplex immunoassay. Male C57BL/6J mice received a high-fat diet plus L-NAME (0.85 g/L) or chow for 5 weeks; after phenotype confirmation, they received oral PF-8380 (30 mg/kg/day) or vehicle for 10 weeks. Endpoints were echocardiography, functional assessment, gravimetric studies, tail-cuff pressure, trichrome fibrosis, and flow cytometry of heart and spleen. Results: All nine analytes, including the autotaxin protein ENPP2, were higher in HFpEF than comparators. HFpEF plasma showed higher LPE O16:1, LPE O18:2, PS 38:4 and PC 36:4;O, and lower SM 39:2; O3 and PS 36:0. LPA 20:0 was 3.5-fold higher in both sexes, whereas LPA 18:2 was lower in women. Diet plus LNAME raised blood pressure, LV mass, and isovolumic relaxation time with preserved ejection fraction. PF-8380 reduced echocardiographic indices of diastolic dysfunction, fibrosis area, cardiomyocyte area, and cardiac CD11b+, CD64+, CD86+, and Ly6G+ frequencies, without altering fat or lean mass. Conclusion: In male mice with established two-hit HFpEF, autotaxin inhibition improved diastolic indices and reduced fibrosis, hypertrophy, and cardiac myeloid accumulation. Human data show altered lysophospholipid composition. Collectively, these findings nominate the autotaxin/LPA axis as a tractable therapeutic target and support further evaluation of autotaxin inhibition as a candidate disease-modifying strategy for HFpEF management.

18
Machine Learning-Based Prediction of Maternal Morbidity across Heterogeneous Populations in the United States using Sequential Modeling of the All of Us Dataset

Zhuang, H.; Zakama, A.; Heller, K.; Faulkner, S.; Gollub, B.; Young-Lin, N.; Chen, I. Y.; Asiedu, M.

2026-08-31 obstetrics and gynecology 10.64898/2026.08.25.26360552 medRxiv
Top 1%
0.5%
Show abstract

In this work, we demonstrate the unprecedented value of NIH's "All of Us Research Program" (AoURP) dataset in studying maternal morbidity and building predictive machine learning (ML) models across heterogeneous populations in the United States. We developed robust and data-driven preprocessing pipelines to curate a longitudinal, multi-site, multimodal, and demographically diverse pregnancy dataset (20,253 subjects; 27,525 pregnancy episodes) from AoURP data, using electronic health records (EHR) (Conditions, Labs, Measurements) and survey responses (Social Determinant of Health (SDoH)), focusing on 7 crucial maternal health adverse outcomes. After characterizing data quality, missingness, and heterogeneity, we performed statistical correlation analysis to identify risk factors. We subsequently developed XGBoost and sequential LSTM models to predict the adverse outcomes, reaching state-of-the-art performance for multiple outcomes. We conducted model interpretability post-hoc analysis to understand success points and fairness analysis to evaluate implications for socio-economic disparities. Four practicing physicians reviewed the set of statistically significant and ML model identified features to assess their clinical validity and novelty. Most features identified through either statistical correlations or ML feature importance analysis aligned with known clinical risk factors. Several features were identified that the ML models used but that are not currently used in clinical practice and may merit further clinical investigation. Fairness analysis revealed certain associations with SDoH and age highlight areas that warrant continued monitoring. Overall, we demonstrate that meaningful populational level patterns can be extracted, and high-performing machine learning models can be trained on this longitudinal, diverse, multi-site dataset. Important risk features, particularly novel ones identified, if validated, could inform new strategies for maternal care or enable development and validation of outcome-specific, clinically deployable ML models.

19
Towards Electronic Health Records-Based Paediatric Growth References: Results from the SwissPedGrowth Project

Leuenberger, L. M.; Shoman, Y.; Romero, F.; Sasaki, M.; Deligianni, X.; Goebel, N.; Mozun, R.; Bielicki, J. A.; Burckhardt, M.-A.; Saner, C.; Schwitzgebel, V.; Hauschild, M.; Righini Grunder, F.; Mueller, P.; Schlapbach, L. J.; Jenni, O.; Spycher, B. D.; Kuehni, C. E.; Belle, F. N.; SwissPedHealth consotrium,

2026-09-02 pediatrics 10.64898/2026.08.28.26361619 medRxiv
Top 1%
0.5%
Show abstract

BACKGROUND: We used anthropometric data from electronic health records (EHRs) of Swiss childrens hospitals to evaluate growth references and estimate centile curves. METHODS: We received EHRs extracted from seven Swiss childrens hospitals and analysed two samples: all children with a height, weight, body mass index (BMI), or head circumference recording, and a subsample restricted to children without diseases potentially affecting growth, weighted to represent the general population. We calculated mean z-scores based on the World Health Organization growth references adopted for Switzerland in 2011 (CH-WHO 2011) and current Swiss growth references (Swiss 2026). We estimated sex-specific centile curves in the subsample using generalised additive models for location, scale, and shape. RESULTS: We included 213,868 children with height, 448,002 with weight, 209,244 with BMI, and 67,397 with head circumference recordings. Mean z-scores in the all children sample were (CH-WHO 2011; Swiss 2026): height (0.10; -0.19), weight (0.16; -0.09), BMI (0.04; -0.07), head circumference (-0.28, -0.28); and in the subsample: height (0.34; 0.00), weight (0.27; 0.01), BMI (0.18; 0.05), and head circumference (0.04; 0.01). The 50th height, weight, BMI, and head circumference centiles of girls and boys in the subsample closely followed those of Swiss 2026, with slightly wider 3rd and 97th centiles in infancy and adolescence. CONCLUSION: Height, weight, BMI, and head circumference centiles aligned well with the Swiss 2026 growth references in Switzerland, demonstrating that hospital EHRs could contribute to future growth references.

20
From Prompt to Provenance: BloClaw, a Capability-Gated AI4S Workstation for Auditable Computational Biology

qin, y.; Pang, J.; Zhang, X.

2026-09-01 bioinformatics 10.64898/2026.08.26.747436 medRxiv
Top 1%
0.5%
Show abstract

Scientific agents can produce plausible answers while remaining unable to establish whether the computation behind an answer is executable, recoverable, or reproducible. We present BloClaw, an AI4S workstation built around a simple principle: a scientific agent should know what it can do, show how it did it, and state what remains unvalidated. Each capability declares an execution state, input constraints, dependencies, expected outputs, and scientific limitations. Natural-language requests are translated into structured tasks, validated against this registry, executed through scientific tools, and recorded in a provenance-aware Living Lab Notebook. The system is designed to detect invalid inputs, failed tool calls, missing dependencies, and remote timeouts, and to route them to repair, retry, or escalation. The implemented and tested scope comprises RDKit-based molecular property and rule screening, protein structure analysis, docking-pose inspection, 3D visualization, and structured reporting. We demonstrate the workflow on a PubChem-retrieved osimertinib structure and a supplied 6LU7 docking artifact: the former yields deterministic descriptors (molecular weight 499.619 Da, cLogP 4.5098, TPSA 87.55 A^2), while the latter contains 2,387 protein ATOM records, 309 residues, and nine pose records. These examples are workflow demonstrations, not efficacy or affinity studies. Beyond retrospective prediction, the manuscript specifies a prior-minimized constructive mode in which a desired function is compiled into explicit physical, chemical, and systems constraints, candidate mechanisms are simulated, and observations are reintroduced for calibration and falsification; this is a proposed extension rather than a result of the present case studies. We describe an evaluation protocol that compares BloClaw with a standard single-agent workflow and fixed-script execution using task completion, scientific correctness, recovery success, provenance completeness, reproducibility, human review time, latency, and cost. This manuscript reports the system design, verified capability boundary, deterministic software artifacts, and a reproducible evaluation protocol; it does not claim benchmark improvements before those experiments are run. BloClaw is an execution and accountability layer for AI-assisted research, complementing expert review and experimental validation rather than replacing them.